Papers with Universal Dependencies project
How Bad are PoS Tagger in Cross-Corpora Settings? Evaluating Annotation Divergence in the UD Project. (N19-1)
Copied to clipboard
| Challenge: | Using annotation variation principles, Part-of-Speech tagging performance degrades when applied to test sentences that depart from training data. |
| Approach: | They propose to use the annotation variation principle to identify inconsistencies between annotations . they also evaluate their impact on prediction performance . |
| Outcome: | The proposed method can detect errors in gold standard annotations and improve prediction performance. |
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)
Copied to clipboard
| Challenge: | a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part. |
| Approach: | They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers. |
| Outcome: | The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets . |
Automatic Extraction of Rules Governing Morphological Agreement (2020.emnlp-main)
Copied to clipboard
Aditi Chaudhary, Antonios Anastasopoulos, Adithya Pratapa, David R. Mortensen, Zaid Sheikh, Yulia Tsvetkov, Graham Neubig
| Challenge: | Creating a descriptive grammar is an indispensable step for language documentation but it is tedious and time-consuming. |
| Approach: | They propose a framework for extracting a first-pass grammatical specification from raw text in a concise, human- and machine-readable format. |
| Outcome: | The proposed framework extracts a grammatical specification that is nearly equivalent to those created with large amounts of gold-standard annotated data. |
Universal Dependencies for Western Sierra Puebla Nahuatl (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated corpus of western Sierra Puebla Nahuatl conforms to universal dependency project annotation guidelines . morphological and syntactic phenomena can be analyzed quantitatively with a large enough corpus . |
| Approach: | They present a morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl . it is the first indigenous language of Mexico to be added to the Universal Dependencies project . UD is a widely-used annotation framework whose aim is to provide a consistent schema for morphological and syntactic phenomena for all of the world's languages. |
| Outcome: | The morpho-syntactically-annotated corpus of western Sierra Puebla Nahuatl conforms to the universal dependency project annotation guidelines. |
Data Augmentation via Dependency Tree Morphing for Low-Resource Languages (D18-1)
Copied to clipboard
| Challenge: | Lack of sizable training datasets leads to poor performance in low-resource languages. |
| Approach: | They propose two techniques to augment training sets of low-resource languages using dependency trees. |
| Outcome: | The proposed methods improve on the training datasets for low-resource languages. |
Errator: a Tool to Help Detect Annotation Errors in the Universal Dependencies Project (L18-1)
Copied to clipboard
| Challenge: | UD project aims to develop cross-linguistically consistent treebank annotations for a wide array of languages. |
| Approach: | They introduce tools that implement the annotation variation principle to help annotators find and correct errors in UD treebanks. |
| Outcome: | The proposed tools can be used to correct errors in UD treebank annotations. |